Papers with task evaluation
On Domain-Adaptive Post-Training for Multimodal Large Language Models (2025.findings-emnlp)
Copied to clipboard
Daixuan Cheng, Shaohan Huang, Ziyu Zhu, Xintong Zhang, Xin Zhao, Zhongzhi Luan, Bo Dai, Zhenliang Zhang
| Challenge: | Adapting general multimodal large language models to specific domains is important for practical applications. |
| Approach: | They investigate domain adaptation of multimodal large language models via post-training . they develop a generate-then-filter pipeline that curates diverse visual instruction tasks . |
| Outcome: | The proposed model outperforms existing models in domain adaptation by combining data from open-source models with training pipelines. |
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows (2024.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) have become a dominant tool for NLP researchers in a wide range of tasks. |
| Approach: | They propose an open source Python library that allows researchers to write simple code to implement powerful LLM workflows. |
| Outcome: | The proposed library is open source and can be used to implement powerful LLM workflows. |
Mining Crowdsourcing Problems from Discussion Forums of Workers (2020.coling-main)
Copied to clipboard
| Challenge: | Among the most widely used platforms are Upwork, Appen, and above all Amazon Mechanical Turk (MTurk) which host annotation tasks and collect huge sets of annotated data from workers. |
| Approach: | They propose to use topic modeling to analyze workers' complaints from a new English corpus of workers’ forum discussions to identify problems in task design, task operation, and task evaluation that workers face with requesters in crowdsourcing processes. |
| Outcome: | The findings form the basis for future research on how to improve crowdsourcing processes. |